Papers with language processing components

2 papers
Domain Expert Platform for Goal-Oriented Dialog Collection (2021.eacl-demos)

Copied to clipboard

Challenge: a prerequisite for the creation of a goal-oriented neural network dialogue system is a dataset that represents typical dialogue scenarios and includes various semantic annotations.
Approach: They propose a web-based platform for collecting and writing goal-oriented dialogue samples.
Outcome: The proposed platform is language-independent and is currently being used to collect dialogue samples in Latvian .
Data Collection Pipeline for Low-Resource Languages: A Case Study on Constructing a Tetun Text Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Labadain Crawler is a data collection pipeline designed to automate and optimize the process of constructing textual corpora from the web, with a specific target to low-resource languages.
Approach: They propose a data collection pipeline built on top of Nutch, an open-source web crawler and data extraction framework, and a tokenizer and identifier for Tetun.
Outcome: The proposed pipeline is based on Nutch, an open-source web crawler and data extraction framework, and is tested with Tetun, one of Timor-Leste’s official languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations